Protein Engineering, Design and Selection
◐ Oxford University Press (OUP)
Preprints posted in the last 90 days, ranked by how well they match Protein Engineering, Design and Selection's content profile, based on 15 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Okuma, A.; Ishida, Y.; Hisada, S.
Show abstract
Chimeric antigen receptor (CAR) T cell therapy has achieved remarkable therapeutic outcomes in hematological cancers. However, broader clinical use has uncovered substantial challenges arising from intrinsic properties of both T cells and tumor tissues. As the functional phenotype of CAR T cells is affected by the CAR molecular architecture, optimizing CAR constructs continues to be a critical and ongoing task. Here, we present a practical workflow for scalable screening of CAR variants in primary T cells using fitness-guided design and mRNA electroporation. Using a CD19-targeted second-generation CAR, we built a library of point mutants that focused mutagenesis on hinge and costimulatory domains. Amino acid substitutions were prioritized using the sequence-based zero-shot fitness predictor to enrich evolutionarily tolerated variants. From 340 designed variants, we electroporated mRNA encoding 85 constructs into primary human CD8+ T cells and quantified cytotoxicity against CD19-positive Nalm6 cells. Twenty-four variants reproducibly exceeded wild-type cytotoxicity across three runs, and three hits were selected for lentiviral validation. One of the selected variants showed significantly improved cytotoxicity despite lower expression frequency and exhibited higher CD62L within CAR-positive cells, suggesting enhanced intrinsic function with a less differentiated phenotype. This approach enables scalable, rapid discovery of improved CAR domain variants directly in primary T cells.
Stark, H.; Faltings, F.; Choi, M.; Xie, Y.; Hur, E.; O'Donnell, T. J.; Bushuiev, A.; Ucar, T.; Passaro, S.; Mao, W.; Reveiz, M.; Bushuiev, R.; Portnoi, T.; Pluskal, T.; Sivic, J.; Kreis, K.; Vahdat, A.; Ray, S.; Goldstein, J. T.; Savinov, A.; Hambalek, J. A.; Gupta, A.; Taquiri-Diaz, D. A.; Zhang, Y.; Snyder, S. J.; Hatstat, A. K.; Arada, A.; Kim, N. H.; Fan, H.; Tackie-Yarboi, E.; Boselli, D.; Schnaider, L.; Liu, C. C.; Li, G.-W.; Hnisz, D.; Sabatini, D. M.; DeGrado, W. F.; Wohlwend, J.; Corso, G.; Barzilay, R.; Jaakkola, T.
Show abstract
We introduce BoltzGen, an all-atom generative model for designing proteins and peptides across all modalities to bind a wide range of biomolecular targets. BoltzGen builds strong structural reasoning capabilities about target-binder interactions into its generative design process. This is achieved by unifying design and structure prediction, resulting in a single model that also reaches state-of-the-art folding performance. BoltzGens generation process can be controlled with a flexible design specification language over covalent bonds, structure constraints, binding sites, and more. We experimentally validate these capabilities in eight diverse design campaigns with functional and affinity readouts across 26 targets. In our experiments, binder modalities span from nanobodies to disulfide-bonded peptides, and targets from disordered proteins to small molecules. In particular, we identify nanobody binders for novel targets with low similarity to proteins with already known bound structures. We release model weights, data, and both inference and training code at: https://github.com/HannesStark/boltzgen.
Zhu, Y.
Show abstract
Antibodies provide programmable molecular recognition, whereas enzymes enable repeated chemical transformation. Catalytic antibodies seek to combine these properties within a single protein scaffold. However, conventional approaches based on transition state analogue immunisation, library screening or local mutagenesis provide limited control over the atomic arrangement of catalytic residues. They also frequently produce antibodies that bind substrates without supporting efficient chemical turnover. Recent advances in generative protein design have enabled the construction of antibody complementarity determining regions and the scaffolding of functional motifs under structural constraints. A systematic strategy for transferring experimentally supported enzyme active site geometry into antibody variable domains is still lacking. Here, we present a computational framework that treats antibody and enzyme structures as distinct but complementary inputs. Developable Fv or VHH structures provide the immunoglobulin scaffold. Enzyme complexes containing substrates, products or transition state analogues provide catalytic residues, ligand conformations, metals, cofactors and key water networks. The selected catalytic atoms are mapped into antibody complementarity determining regions, while the surrounding loops are reconstructed using antibody compatible representations and constrained all atom diffusion. Sequence design and structural back prediction are followed by filters for antibody folding, catalytic geometry, ligand positioning, conformational stability and developability. The framework avoids direct fusion of intact enzymes and antibodies. Instead, it transfers only the local geometry required for catalysis. This separation of scaffold selection from catalytic motif selection creates a testable route for determining whether natural enzyme chemistry can be embedded within antibody formats. It also provides a practical basis for evaluating substrate binding, chemical conversion, product release and catalytic turnover as separate design objectives. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=91 SRC="FIGDIR/small/740676v1_ufig1.gif" ALT="Figure 1"> View larger version (45K): org.highwire.dtl.DTLVardef@16ef6d6org.highwire.dtl.DTLVardef@f7c23org.highwire.dtl.DTLVardef@9ee50borg.highwire.dtl.DTLVardef@1cf42d4_HPS_FORMAT_FIGEXP M_FIG C_FIG
Bozkurt, C.; Nathanail, E.; Goteti, A.
Show abstract
For structural-biology and protein-production pipelines, the hardest part of a difficult protein is not the biology -- it is obtaining a well-behaved sample for functional studies. Programs routinely stall at construct design, expression, and purification: deciding where to truncate, which tags to use, how to express, and how to purify so the protein survives concentration and handling. These decisions are still made largely by literature precedent and experimental experience, and they require trial-and-error before arriving at a functional construct for hard targets. We present a prospective, single-pair wet-lab case study testing whether an integrated computational platform can improve these decisions. For human fibroblast growth factor 21 (FGF21) -- a clinically important and stability-challenged metabolic hormone -- we compared two expression constructs produced side by side under the same experimental workflow, using two different design strategies: one designed by a scientist from the literature (reproducing the published core-domain construct, PDB 6M6E), and one designed by the Orbion platform -- an AI, prediction-guided protein-design system (orbion.life) -- which additionally generated the expression and purification protocols (executed scientist-in-the-loop). The platforms construct used an unconventional, longer C-terminal boundary not found in public sequence databases. Since the two constructs differ in more than one feature, we treat them as workflow-level designs throughout. The scientist construct gave a higher initial yield ([~]2.4 xmore protein recovered at affinity capture). The platform-designed construct, however, showed a more favourable downstream developability profile: it concentrated higher (1.4 vs 0.7 mg/mL) while remaining more monodisperse by dynamic light scattering (DLS). The scientist construct, in contrast, aggregated on concentration, so its initial-yield advantage did not survive: in the final concentrated sample the Orbion construct provided the more usable material for downstream studies. Computed for the mammalian host used, the platform had prospectively scored its own design higher (composite 68.7 vs 59.0 for the scientist-designed construct), and its predictions of yield, solubility, and disorder matched the wet-lab outcome. This is a single, deliberately scoped case study, not a population-level benchmark; the two constructs differ in more than one feature, and biological activity was not assayed. Alongside the bottlenecks of this approach discussed here, used as a decision aid, prediction-guided construct and protocol design has the potential to remove costly iteration cycles of protein production campaigns.
Mavar, L.; Pavlenok, M.; Paul, A.; Hall, L.; Larimer, B. M.; Niederweis, M.
Show abstract
Calreticulin is an emerging cancer biomarker, but current detection methods rely on expensive monoclonal antibodies that suffer from inefficient protein production, pharmacokinetic challenges and poor tissue penetration. Cal3, a calreticulin-specific nanobody, was constructed by replacing the complimentary determining region 2 (CDR2) of a soluble, clinically validated nanobody with a calreticulin-specific CDR2 isolated from a phage display library. However, the poor solubility and low yield of Cal3 limit its usefulness. In this study, we engineered CALR-Nb02 by adapting the core of Cal3 to a partial consensus framework sequence of stable nanobodies. CALR-Nb02 was purified with a 240-fold higher yield as a predominantly monomeric, soluble protein that exhibits an increased thermal stability and a higher calreticulin binding affinity (KD: 25-50 nM) compared with Cal3. These results reveal a strategy for quickly altering the specificity of a stable nanobody, and provide an improved calreticulin-binding reagent for future diagnostic, imaging, and therapeutic applications.
Kim, Y.; Kwon, H.; Song, J.; Lee, Y.; Park, M.; Lee, C.-H.
Show abstract
Therapeutic antibody development requires workflows that integrate antigen-reactive clone discovery with efficient humanization and early developability assessment. Here, we combined immune yeast fragment antigen-binding (Fab) display with single-round focused humanization and applied the workflow to antibodies against amyloid-{beta} (A{beta})-derived preparations. Immunization with A{beta}1-42 aggregate preparations generated a Fab-display library with a diversity of approximately 3.5 x 108. Magnetic enrichment followed by fluorescence-activated cell sorting (FACS) identified three sequence-distinct immunoglobulin G (IgG)-format candidates, of which CLAB17 and CLAB45 were advanced to humanization. Structure-guided libraries sampled framework positions predicted to support complementarity-determining regions (CDRs) or heavy-and light-chain variable-domain packing, and a single FACS round recovered binding-positive variants CLAB17-h2 and CLAB45-h8. Both retained the parental CDRs and showed increased predicted humanness, favorable computational developability triage profiles, and high purity by sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE). By enzyme-linked immunosorbent assay (ELISA), CLAB17-h2 showed a lower apparent half-maximal effective concentration (EC50) for A{beta}1-42AggreSure, whereas CLAB45-h8 showed a lower apparent EC50 for pyroglutamate-modified A{beta}3-42 (A{beta}pE3-42). Because the preparations were not resolved into defined assembly states, these antibodies are considered A{beta}-preparation-binding rather than aggregate-state-selective candidates. This workflow provides a practical route from immune-repertoire discovery to binding-positive humanized antibodies.
Rawat, P.; Kyte, J. A.; Greiff, V.; Dorraji, E.
Show abstract
Human epidermal growth factor receptor 2 (HER2) is an oncogenic receptor tyrosine kinase in breast cancer and other malignancies. A subset of HER2-positive tumours expresses 611-CTF-p95HER2, a tumour-specific, hyperactive truncated isoform associated with metastasis and treatment resistance that lacks most of the extracellular domain targeted by conventional HER2-directed antibodies. We previously developed NAZ-mAb (formerly known as Oslo-2), a monoclonal antibody against 611-CTF-p95HER2. Here, we describe a computational antibody-engineering workflow for designing variants of NAZ-mAb. Starting from the sequence alone, we modeled the NAZ-mAb-611-CTF-p95HER2 complex, generated a combinatorial mutational landscape using FoldX 5.0, and prioritized candidate variants using predicted interaction energy and developability criteria. Two variants representing distinct design strategies were selected for validation: an aromatic double mutant, NAZ-mAb v1 (L:S31W/L:H107W), and a conservative single mutant, NAZ-mAb v2 (L:S31M). Both variants were successfully expressed as recombinant IgGs; NAZ-mAb v2 achieved a five-fold higher recombinant expression yield than parental NAZ-mAb, while both variants retained antigen binding with a higher apparent signal than the parental antibody in indirect ELISA. However, Biacore two-state kinetic analysis revealed weaker affinities than the parental antibody (KD NAZ-mAb v1: 32.6 nM, NAZ-mAb v2: 9.45 nM vs. parental NAZ-mAb: 5.33 nM). These findings show that the computational workflow can generate experimentally tractable, antigen-engaging NAZ-mAb variants, while also highlighting the limitations of fixed-backbone interaction-energy ranking as a predictor of binding affinity and yield. This study provides a practical framework for computationally driven, developability-aware antibody optimization in the absence of experimental structural data.
Stephenson, H.; Voicu, D.; Novakov, V.; Levy, M.; Marsilio, J.
Show abstract
With the growing use of machine-learning-assisted pipelines for designing, characterizing, and optimizing biomolecules, the reliability of structure prediction models is increasingly important. PolyFold is a benchmarking framework developed to evaluate open-use structure prediction models, Boltz-2 and OpenFold 3, as commercially accessible alternatives to AlphaFold 3. We outline an end-to-end workflow automation tool to streamline input file creation, batch automation, and comprehensive analysis of model outputs for leading open-use structure prediction models. We curated an evaluation dataset of several thousand high-quality Protein Data Bank structures, homology-filtering against the training sets of both models to ensure a fair analysis. We then implemented an evaluation pipeline incorporating structural metrics (RMSD, TM-score, lDDT, etc.), interface metrics (DockQ, ilDDT, iRMSD, etc.), and physicochemical realism checks (based on bond lengths, angles, molecular internal energies, etc.). We identify key performance disparities, observing that Boltz-2 is generally superior to OpenFold 3, though the differential is partially attributable to residual homology leakage not accounted for by prevailing test set curation practices. We thus recommend a new method for homology-reducing when building a test set using length-weighted average fractional identity cutoffs rather than lowest chain fractional identity cutoffs. Even in eliminating residual leakage, Boltz-2 still performs better on full-set comparisons and a variety of important partitions (nucleic acids, protein-ligands, Ab-Ags, etc.). Both models are strong at folding monomeric structures, though struggle with homomultimer placement and small molecule physical realism, demonstrating enduring limitations of machine learning methods. This work is the first end-to-end, open-use, and reproducible platform for systematically assessing state-of-the-art structure prediction models. PolyFold enables practitioners to determine how models compare in performance on specific inference tasks and supports the broader adoption of accessible computational tools to facilitate biomolecular science.
Lee, S.; Tak, E.-J.; Shim, H.-J.; Ahn, W.-C.; Park, K.-H.; Go, S.-R.; Yang, H.; Woo, E.-J.
Show abstract
Flavin adenine dinucleotide-dependent glucose dehydrogenase (FAD-GDH) is a redox enzyme widely used in glucose monitoring, bioelectronic devices, and enzymatic biofuel cells because of its oxygen-independent catalysis and compatibility with electron-transfer processes. However, protein-based regulators that directly bind GDH and modulate its redox output remain underdeveloped. Here, we present an AI-guided strategy for developing a de novo protein inhibitor targeting FAD-GDH. GDH-targeting candidates generated through structure-based computational design were evaluated by yeast surface display and fluorescence-activated cell sorting, leading to the identification of FAD-GDH inhibitor-1 (FGI-1) as a GDH-targeting inhibitory scaffold. Purified His-MBP-FGI-1 reduced GDH-mediated DCIP reduction, demonstrating attenuation of GDH-derived redox output. Random mutagenesis followed by secondary FACS screening yielded evolved variants with increased GDH-binding signals and enhanced redox-output suppression, showing that the de novo inhibitory scaffold could be functionally tuned through experimental evolution. In addition, an FGI-1-based construct fused to a larger protein module retained GDH-output suppressive activity, and electrode-based measurements showed reduced GDH-derived current output. Because electrode-associated measurements may be influenced by protein-mediated surface shielding and altered electron-transfer accessibility, this decrease was interpreted conservatively as attenuation of GDH-derived electrochemical output rather than direct evidence of active-site inhibition. Together, this work establishes an AI-guided design-validation workflow for developing protein inhibitors that modulate FAD-GDH redox output and provides a foundation for protein-level control of enzyme output in biosensing and bioelectronic applications.
Liu, D.; Sreenivasan, S.; Gray, C. J.; Cleveland, H. C.; Swint-Kruse, L.
Show abstract
A central challenge in molecular biology is understanding how amino acid substitutions modulate various features of protein function and stability. To illuminate the complexities of this relationship, high-throughput (HTP) assays are increasingly used to assess site-saturating mutagenesis libraries. A common downstream analysis is to average the set of twenty outcomes at each amino acid position for comparison with structural and evolutionary features. Average values clearly identify positions that tolerate most substitutions (neutral positions) and positions where most substitutions abolish activity (toggle positions). However, average values conceal the existence of rheostat positions, where different amino acid substitutions sample a wide range of outcomes. To quantitatively identify rheostat positions, we previously developed a histogram-based analysis that we here expand by: (i) incorporating new position classes observed in experimental studies of rheostat positions; (ii) formalizing a hierarchy of class assignments; (iii) refining error-based identification of neutral positions; and (iv) statistically assessing the robustness of class assignments to changes in experimental and computational parameters. RheoScale 2.0 is implemented in Excel and newly implemented in Python for facile integration with existing HTP pipelines; all parameters are customizable. Example analyses are shown for three HTP datasets of the SARS-CoV-2 papain-like protease. Results illustrate two aspects that influence interpretation of HTP data: First, position assignments (and substitution outcomes) depend highly on the measured feature. Second, many protein positions play multiple roles in the sequence-structure-function relationship. The recognition of varied position roles will advance understanding of pathogen evolution, protein engineering, and variant interpretation for personalized medicine. SummaryRheoScale 2.0 improves how high-throughput mutational data are interpreted by identifying protein positions where amino acid substitutions act like biological dimmer switches. By enabling more nuanced assignment of position behavior, beyond neutral or deleterious outcomes, this analysis framework advances studies of sequence-structure-function relationships and has broad relevance for understanding protein evolution, engineering proteins with desired properties, and interpreting variants linked to human disease. SOFTWARE AVAILABILITYhttps://github.com/liskinsk/RheoScale-calculator
Lee, A. L.; Seamann, A.; Chung, G.; Zhao, C.; Maddamsetti, R.; Khare, S.
Show abstract
Machine learning-guided protein sequence redesign is now routinely used to optimize multiple properties relevant for protein engineering, most prominently thermostability and recombinant expression levels. Salt tolerance is a valuable property for "blue-biotechnology"- enabled biomanufacturing, yet no generative computational method exists to redesign proteins for increased salt tolerance. We hypothesized that a training dataset heavily biased toward salt-adapted proteomes would yield a model capable of designing proteins with halophilic properties. To test this, we retrained the sequence redesign model ProteinMPNN on proteins from "salt-in" extreme halophiles such as Haloarcula marismortui, a Dead Sea archaeon that grows optimally near 3-4 M NaCl, roughly six times the salinity of seawater, and accumulates molar concentrations of salts in its cytoplasm. Our model, HaloMPNN, redesigns non-halophilic proteins so that their properties shift towards those of natural halophilic proteins: lower predicted isoelectric point, greater surface acidity, and reduced surface and core hydrophobicity. Redesigning a broad range of non-halophilic proteins with SolubleMPNN, ProteinMPNN, and HyperMPNN shows that this shift is specific to HaloMPNN rather than a generic consequence of sequence redesign. HaloMPNN therefore offers both a route to designing candidate salt-tolerant enzymes and a means of identifying the characteristics that underlie halophilic adaptation.
Moranzoni, G.; Jorgensen, L. V.; del Cerro, J. H.; Andreoletti, A.; Hoie, M. H.; Vitting-Seerup, K.; Barnkob, M. B.; Olsen, L. R.
Show abstract
Chimeric antigen receptor (CAR) cell therapy has achieved transformative clinical success through targeting of CD19 in refractory B cell malignancies, but extension of this strategy to solid tumors, other hematological malignancies, and autoimmune disease has exposed the complexity of target selection. Antigen abundance alone is not sufficient to define a suitable CAR target. Instead, therapeutic efficacy and safety are shaped by a broader set of molecular features, including isoform usage, subcellular localization, secretion, epitope stability, and the structural context in which antibody-derived binding domains engage their target. At the same time, advances in transcriptomics, structural biology, and artificial intelligence (AI)-enabled prediction now make it possible to assess many of these properties systematically. Here, we outline the principal molecular features that characterize effective and safe CAR targets and present a practical framework that integrates public datasets with computational and AI-based tools for their evaluation. Using HER2 as an illustrative case, we show how isoform-resolved expression, single-cell analyses, topology prediction, structure modelling, epitope mapping, and in silico binding analyses can reveal liabilities that are not captured by conventional target-expression screens alone. This framework provides a systematic strategy to prioritize targets and epitopes, guide preclinical investigation, and de-risk clinical translation. We anticipate that such integrative workflows will become increasingly important for moving CAR target discovery from descriptive expression analysis towards informed therapeutic design.
Styles, M. J.; Xie, V. C.; Legault, S.; Pixley, J. A.; Dickinson, B. C.
Show abstract
Aberrant protein-protein interactions (PPIs) drive myriad diseases. Inhibiting these PPIs often relies on discovering molecules that bind to one of the proteins and hoping that this binding inhibits the PPI. Molecular binder discovery often takes months, but a discovery process that ensures that the resulting molecule not only binds a target protein, but selectively inhibits a target PPI, could dramatically accelerate these endeavors. Here, we develop Phage-Assisted Non-Continuous Selection of PPI Inhibitors (PANCS-Inhibitors): a rapid screening platform that directly selects for molecules capable of disrupting a pre-formed PPI. We demonstrate this new platform using three clinically relevant oncogenic PPIs: KRas-Raf, Mdm2-p53, and Myc-Max. PANCS-Inhibitors can be used to both improve known PPI inhibitors and for de novo discovery of mini-protein PPI inhibitors that function in mammalian cells. This platform has the potential to rapidly generate inhibitors for many clinically relevant PPIs, which can be used as starting points for therapeutic development.
Bayat, P.; Perkins, S. J.; Clancy, S.; Patel, S. S.; Yin, R. F.; Bozovicar, K.; Singh, S.; Shrestha, S.; Moustafa, Z.; Zayani, R.; IWE, I.; Bayat, S.; Kelly, P.; Vigar, J. R. J.; White, V. Y.; Xie, M.; Simchi, M.; Palter, S.; Nguyen, J.; Zeisler, I. Y.; Wu, B.; Pardee, K.
Show abstract
Discovering functional peptides across vast sequence space remains a formidable challenge, particularly when experimental training data is scarce. We present Minimal Data Maximal Insight (MDMI), a two-stage structure-guided computational pipeline that designs functional peptide variants using only a small, annotated dataset. Rather than relying on sequence information alone, MDMI integrates three-dimensional structural features derived from predicted peptide-protein complexes into a machine learning model that captures interface geometry and binding energetics. This structure-aware predictor, paired with a genetic algorithm for sequence exploration, reduced false positives from 70% to close to zero in an all-negative benchmark panel compared with a sequence-only model in computational benchmarking, and produced approximately four-fold more high-confidence in silico binders than state-of-the-art peptide/protein design baselines. Using the split-GFP system as a testbed, where fluorescence provides a direct functional readout of peptide-protein complementation, MDMI identified peptides with up to 38% sequence divergence from wild-type in Stage 1 while retaining measurable activity. In Stage 2, motif-guided recombination of successful Stage 1 variants produced highly divergent yet functional peptides bearing over 50% sequence difference from wild-type, revealing two distinct functional clusters in sequence space. As further validation, a top-performing candidate expressed as a full-length GFP fusion retained a GFP-like emission profile, supporting formation of a fluorescent GFP-like scaffold. These results demonstrate that structure-informed pipelines can uncover remote functional sequence space from minimal data, with broad implications for peptide and therapeutic analog discovery.
Gutierrez, D.; Madrigal Harrison, I.; Feller, A.; Ellington, A.
Show abstract
L-3,4-dihydroxyphenylalanine (L-Dopa) is an important pharmaceutical for the treatment of Parkinsons disease and a precursor to numerous catechol-containing compounds. The flavin-dependent monooxygenase HpaBC is a promising biocatalyst for microbial L-Dopa production but exhibits limited native activity toward L-tyrosine. Although structure-based machine learning (ML) models have become increasingly popular for protein engineering, relatively few studies have systematically compared their performance or evaluated their integration into iterative engineering workflows. Here, we benchmarked multiple ML models for their ability to predict activity enhancing mutations in HpaBC. Experimentally validated single mutants were used to seed combinatorial design with EVOLVEpro, generating progressively improved higher-order variants. We next evaluated how expanding the EVOLVEpro training set with directed evolution derived variants influenced combinatorial predictions and finally explored an expanded sequence space by allowing combinations of both machine learning derived and directed evolution derived mutations. This workflow produced HpaBC variants with substantially improved activity. Although incorporating directed evolution data substantially altered EVOLVEpros predicted mutational trajectories, both training strategies converged on variants with comparable activities, demonstrating that distinct regions of sequence space can yield similarly optimized enzymes. Together, these results provide a systematic comparison of zero-shot ML models and establish an iterative framework for integrating machine learning with directed evolution to accelerate enzyme engineering.
Hu, A.; Bailey, J. S.; Spina, S. C.; Rajagopal, G.; Phan, N.; Kimmel, B. R.
Show abstract
Targeted covalent inhibitors are a powerful, yet underexplored, class of therapeutics, and current computational covalent screeners are constrained in early drug discovery due to the need for prior knowledge of the target site and limited throughput. We present CovSite, a blind covalent screening tool that identifies candidate reactive residues across the entire protein surface, utilizing only the protein structure and electrophile SMILES. CovSite applies a pipeline of four orthogonal physicochemical filters (nucleophile identification, solvent accessibility, environment-dependent deprotonation prediction, and semi-quantum-mechanical reactivity ranking) to identify potential small-molecule candidate inhibitors. Validated against 2,062 diverse covalent protein-ligand complexes spanning six nucleophilic residue types, CovSite achieves a 98.5% blind target site hit on a held-out benchmark set of 207 cysteine-targeted complexes while reducing the search space by 97.8%. The target-site hit detection exceeds the 53-62% accuracy of popular covalent screening tools operating under non-blind conditions on the same benchmark set. By extending nucleophilic coverage beyond cysteine to include serine, threonine, lysine, histidine, and tyrosine, and completing a screening of a 200-residue protein in two to three minutes on standard hardware, CovSite serves as a platform technology with the potential to address critical gaps in throughput, generalizability, and accuracy in this field of covalent screening. We demonstrate this capability by using CovSite as a blind, ligand-specific approach that enables iterative, machine-learning-driven covalent inhibitor generation that is impractical with existing tools, establishing a foundation for computationally guided covalent drug discovery for novel and understudied targets. TOC Figure O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=100 SRC="FIGDIR/small/739288v1_ufig1.gif" ALT="Figure 1"> View larger version (27K): org.highwire.dtl.DTLVardef@17abeeforg.highwire.dtl.DTLVardef@18d6299org.highwire.dtl.DTLVardef@14437eaorg.highwire.dtl.DTLVardef@1b2ec81_HPS_FORMAT_FIGEXP M_FIG C_FIG
Li, M.; Cheng, X.; Jiang, F.; Hong, L.; Yu, Y.
Show abstract
Designing mutations that enhance protein stability is a central goal in protein engineering. However, experimentally screening large numbers of candidate mutations is costly and time-consuming, creating a strong need for computational methods that can identify potentially stabilizing mutations. Among these approaches, protein language models are particularly promising because they learn context-dependent amino acid preferences from large-scale sequence and structure datasets. Nevertheless, most existing stability prediction methods use these models primarily as feature extractors and do not fully exploit the amino acid probability distributions they encode. Here, we introduce MAXWELL (Matrix-wise Landscape Learning), a novel post-training method that calibrates the probabilistic outputs learned by protein language models during pretraining to generate mutational landscapes that quantify the effects of individual amino acid substitutions on protein stability. When applied to ProteinMPNN, MAXWELL yields a state-of-the-art predictor of the effects of protein mutations on stability, outperforming ThermoMPNN and other representative methods on a curated benchmark of experimentally measured stability changes. We next applied MAXWELL to the design of ten single-point mutations in the DhaA dehalogenase, seven of which (70%) increased thermal stability. Among them, G171W showed the largest improvement, with a measured {Delta}Tm of 4.91 {degrees}C. These experimental results establish MAXWELL as a novel post-training strategy for protein language models and a practical framework for designing stabilizing mutations. Repositoryhttps://github.com/ai4protein/Venus-MAXWELL
Sharma, S.; Ramachandran, V.; Komath, S. S.; Muthuswami, R.; Gourinath, S.
Show abstract
Epigenetic regulation of chromatin dynamics via histone acetylation is one of several mechanisms by which eukaryotes regulate gene expression, DNA replication and repair, and maintain genome stability. This function is performed by histone acetyltransferases (HATs). Rtt109 is one such cytoplasmically localized HAT required for H3K56 acetylation found exclusively in fungi. Using recombinantly expressed Candida albicans Rtt109 and its chaperones, Vps75 and Asf1, we show that it can acetylate a 20-residue N-terminal H3 peptide in a coupled HAT assay only in the presence of Vps75, but not in the presence of Asf1 in vitro. This appears to be due to the fact that Rtt109-Vps75 is a high affinity stable complex, as estimated by biolayer interferometry (BLI) and gel filtration studies. The HAT activity of the Rtt109-Vps75 complex necessarily requires a flexible 118-160 residue loop of Rtt109 but not the C-terminal domain of Vps75. These results are comparable with what has been observed for the Saccharomyces cerevisiae Rtt109 homolog. In silico screening of 1,350,000 molecules from Life Chemicals Databases identified some likely inhibitors of C. albicans Rtt109 and six of them tested for binding to Rtt109 using BLI. The best ligand, F2368-0266, was used to study its effect on steady state enzyme kinetics, and found to be a competitive inhibitor of the peptide substrate but not of acetyl-CoA. Given the importance of Rtt109 in regulating virulence attributes such as hyphal morphogenesis and GPI biosynthesis in Candida albicans, and its effect on fungal pathogenesis, these results have significant clinical implications.
Alejo, K.; Korban, C.; Chung, C.
Show abstract
Structure-based drug discovery is known to apply computational methods in a tiered hierarchy, with each layer narrowing the candidate set and refining the binding picture before committing to the next, more expensive step. We present a four-tiered computational benchmarking study evaluating five engines against a panel of 36 compounds targeting B-secretase 1 (BACE1), a validated Alzheimer's disease target with extensive co-crystal ground truth. This study evaluates Flexible Docking and Boltz2 Cofolding as the primary tier, followed by Ensemble Docking, and then Protein-Ligand MD with MM/PBSA and MM/GBSA post-processing. This is then concluded with Relative Binding Free Energy Perturbation (RevFEP) as the terminal refinement layer. Each method was benchmarked against the experimental binding free energies derived from the co-crystal structures spanning -7.85 to -11.35 kcal/mol. Our findings revealed that Flexible Docking reproduced the co-crystal binding mode for 35 of 36 ligands (97.2% within 2.0 A RMSD) but did not rank potency at this resolution. Boltz2 CoFolding provided an orthogonal structural cross-check with a receptor backbone RMSD of 0.293 A against the experimental co-crystal structure. Ensemble Docking identified the optimal receptor conformation for downstream FEP setup. MD with MM/GBSA decomposition identified van der Waals complementarity as the primary potency driver (Pearson r = +0.855, R2 = 0.732 on a 10-compound subset). RevFEP delivered the highest affinity correlation of any method (Pearson r = +0.662, R2 = 0.438, Spearman p = +0.624, mean absolute error 1.02 kcal/mol across all 36 ligands), resolving potency differences within a narrow 3.5 kcal/mol congeneric window that no other engine could discriminate. We characterize what each engine contributes independently and where RevFEP delivers signals no other engine achieves.
Ucar, T.; Bates, J.; Fu, Y.; Shi, W.; Stark, H.; Nava, D.; Cavalleri, L.; Wohlwend, J.; Corso, G.; Passaro, S.
Show abstract
Designing binders against novel protein targets remains a central challenge in computational drug discovery. Here we introduce BoltzProt-1, a pipeline for generating protein binders, including nanobodies, with improved hit rates and favorable developability properties. At its core lie a refined iteration of BoltzGens generative model and a novel protein-protein interaction prediction model, BoltzPPI. Employing BoltzPPI instead of BoltzGens standard structure-prediction confidence metrics to rank nanobody (VHH) designs increases the confirmed-binder hit rate from 3.3% to 8.0% across 10 novel targets. Assessed on 10 additional targets used in prior literature, the BoltzProt-1 pipeline obtains nanobody screening hits for 7 of 10 targets, surpassing the 6 of 10 previously reported by Chai-2. Finally, evaluating the developability of BoltzProt-1-designed nanobodies in terms of stability, aggregation, purity, polyspecificity and hydrophobicity reveals that 58% of its confirmed binders pass every criterion, exceeding both BoltzGen (40%) and clinical-stage VHH controls (21%). O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=104 SRC="FIGDIR/small/733997v1_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@125fb31org.highwire.dtl.DTLVardef@8e7482org.highwire.dtl.DTLVardef@8318a1org.highwire.dtl.DTLVardef@c62ab5_HPS_FORMAT_FIGEXP M_FIG C_FIG